Asymptotic Behavior of Multivariate Reward Processes with Nonlinear Reward Functions

Authors

A. R. Soltani

K. Khorshidian

Abstract:

This article doesn't have abstract

Download for Free

Already have an account?login

similar resources

asymptotic behavior of multivariate reward processes with nonlinear reward functions

full text

COVARIANCE MATRIX OF MULTIVARIATE REWARD PROCESSES WITH NONLINEAR REWARD FUNCTIONS

Multivariate reward processes with reward functions of constant rates, defined on a semi-Markov process, first were studied by Masuda and Sumita, 1991. Reward processes with nonlinear reward functions were introduced in Soltani, 1996. In this work we study a multivariate process , , where are reward processes with nonlinear reward functions respectively. The Laplace transform of the covar...

full text

On Markov Decision Processes with Pseudo-Boolean Reward Functions

full text

Asymptotics for renewal-reward processes with retrospective reward structure

Let {(Xi; Yi): i= : : : ;−1; 0; 1; : : :} be a doubly in nite renewal-reward process, where {Xi: i= : : :− 1; 0; 1; : : :} is an i.i.d. sequence of renewal cycle lengths and Yi= g(Xi−q; Xi−q+1; : : : ; Xi) is the lump reward earned at the end of the ith renewal cycle, with some function g :R q+1 → R . Starting with the rst renewal cycle (of duration X1) at the time origin, let C(t) denote the e...

full text

Markov Decision Processes with Arbitrary Reward Processes

We consider a learning problem where the decision maker interacts with a standard Markov decision process, with the exception that the reward functions vary arbitrarily over time. We show that, against every possible realization of the reward process, the agent can perform as well—in hindsight—as every stationary policy. This generalizes the classical no-regret result for repeated games. Specif...

full text

Sparse Reward Processes

We introduce a class of learning problems where the agent is presented with a series of tasks. Intuitively, if there is a relation among those tasks, then the information gained during execution of one task has value for the execution of another task. Consequently, the agent is intrinsically motivated to explore its environment beyond the degree necessary to solve the current task it has at han...

full text

My Resources

Save resource for easier access later

Save to my library Already added to my library

{@ msg_add @}

Journal title

Bulletin of the Iranian Mathematical Society

volume 28 issue No. 2

pages 1- 17

publication date 2011-01-24

unfollow

{@ msg @}

By following a journal you will be notified via email when a new issue of this journal is published.

Keywords

Semi-Markov processes Reward processes Laplace transform

Hosted on Doprax cloud platform doprax.com